Conversation
3 tasks done
zzylol
force-pushed
the
stack/univmon-l2-accuracy
branch
from
October 5, 2026 03:07
18cb590 to
ba716f2
Compare
zzylol
changed the base branch from
stack/509-viewer-stages
to
stack/editor-sql
October 5, 2026 03:08
This was referenced Oct 5, 2026
UnivMon's L2 readout now reads layer 0's CountSketch row-median F2 (the textbook AMS estimator over the whole stream) instead of the heavy-hitter G-sum heuristic, and the built-in model certifies it: Chebyshev per row (p = 0.1, eps2 = sqrt(20/w)), binomial median tail over odd d rows, relative L2 bound 1 - sqrt(1 - eps2). Fails closed unless d is odd, w a power of two and d*log2(w) + d <= 128. UnivMon is sized by inverting that bound for every readout, and Stage 3 prices it from its rows and Pass 1's state-size function. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#620's priced_example2_candidates_bind requires every priced Stage 3 candidate to bind. With UnivMon L2 certified, Example 2's UnivMon over src_ip (unit-weight item update) is now priced, but the physical planner sent it to the keyed weighted-frequency build, which has no UnivMon kernel. Bind it as UnivMon's value-frequency build over the item column, which already counts typed SQL identities. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
zzylol
force-pushed
the
stack/univmon-l2-accuracy
branch
from
October 5, 2026 06:21
ba716f2 to
e34c85c
Compare
This was referenced Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stack: #574 → #620 → #618 → #621 → #627 → #625 → #628 → #632 → #634 → #616 → #617 → #622 → #624 → #629 → #630 → #631 → #633 → #635 → #636 → #637
Problem
Stage 3 rejected every UnivMon candidate ("no accuracy model for UnivMon"). The only certified readout was the exact total. The kernel read L2 via
calc_l2, the square root of the heavy-hitter G-sum heuristic, and no theory covers that readout. UnivMon's shape was a fixed 5×1024×16 whatever the requirement, and Stage 3 priced it through the catch-all(1, 1_024).Changes
FrequencyL2now readsl2_sketch_layers[0].get_l2(). That is √(median over rows of Σ C²) on layer 0, which sees the whole stream. It is the textbook AMS F₂ estimator and is linear under merge. No sketchlib change.estimators/univmon.rs):FrequencyL2is now certified.RelativeValueunder contractunivmon_layer0_f2_median_chebyshev_v1. The contract records the idealized 128-bit-hash assumption, p and the readout.None) unless d is odd, w is a power of two and d·log₂w + d ≤ 128. sketchlib slices each row's column and sign from one 128-bit hash, so outside these conditions the rows are not independent.summary_shapegets a UnivMon arm. Ops per row are 2·(2d + 4), since an insert reaches about 2 layers. State bytes come from Pass 1'ssketch_state_bytes.RelativeValue; even d, non-power-of-two w and bit-budget overflow giveNone; tighter ε/δ give more columns/rows; F0 and entropy stayNone.get_l2, and on a Zipf(1) stream at the (0.01, 0.01) sizing it is within 1% of the true L2.build_violation, while Q1 and Q2 still fail with "no accuracy model".frontend-promql/tests/univmon_candidates.rs: count now uses the same ε as the other three readouts, so all four share one shape. L2 is now expected to be certified, and the "uncalibrated" test covers entropy only.integration-tests/tests/precompute_raw_samples.rs: raises this test's executormax_bytesto 1 GiB. A UnivMon sized at ε = 0.02 is about 10 MiB per population (16 layers × 5 × 2¹⁴ × 8 B), so five series exceed the default 64 MiB.planner/tests/summary_sharing.rs(stage-pipeline sharing test): with synthetic evidence, the plan whose three estimates read one UnivMon build is valid. It is no longer selected, because the new pricing (28 updates/row) costs more than exact counting. The test now asserts that the valid shared candidate exists.accuracy-models.mdtable row andunivmon-frequency-summary.mdupdated. Doc comments updated indevtools/tests/stage_pipeline.rsandsummary_sharing.rs.Example 2 (
stage_pipeline --example planner-layering-2):tools/dag-viewer/examples/planner-layering-example1.jsonis unchanged (the fixture test passes).src_ip(a unit-weight item update) is now priced, but the executor's physical planner sent it to the keyed weighted-frequency build, which has no UnivMon kernel. It now binds as UnivMon's value-frequency build over the item column, which already counts typed SQL identities (priced_example2_candidates_bind).Test plan
cargo fmt --all --checkcargo clippy --workspace --all-targets -- -D warningscargo test --workspace: 1677 passed, 0 failed, 21 ignored (re-run after the rebase on main d4869a7)🤖 Generated with Claude Code